Drug-likeness analysis of traditional Chinese medicines: prediction of drug-likeness using machine learning approaches.
نویسندگان
چکیده
Quantitative or qualitative characterization of the drug-like features of known drugs may help medicinal and computational chemists to select higher quality drug leads from a huge pool of compounds and to improve the efficiency of drug design pipelines. For this purpose, the theoretical models for drug-likeness to discriminate between drug-like and non-drug-like based on molecular physicochemical properties and structural fingerprints were developed by using the naive Bayesian classification (NBC) and recursive partitioning (RP) techniques, and then the drug-likeness of the compounds from the Traditional Chinese Medicine Compound Database (TCMCD) was evaluated. First, the impact of molecular physicochemical properties and structural fingerprints on the prediction accuracy of drug-likeness was examined. We found that, compared with simple molecular properties, structural fingerprints were more essential for the accurate prediction of drug-likeness. Then, a variety of Bayesian classifiers were constructed by changing the ratio of drug-like to non-drug-like molecules and the size of the training set. The results indicate that the prediction accuracy of the Bayesian classifiers was closely related to the size and the degree of the balance of the training set. When a balanced training set was used, the best Bayesian classifier based on 21 physicochemical properties and the LCFP_6 fingerprint set yielded an overall leave-one-out (LOO) cross-validated accuracy of 91.4% for the 140,000 molecules in the training set and 90.9% for the 40,000 molecules in the test set. In addition, the RP classifiers with different maximum depth were constructed and compared with the Bayesian classifiers, and we found that the best Bayesian classifier outperformed the best RP model with respect to overall prediction accuracy. Moreover, the Bayesian classifier employing structural fingerprints highlights the important substructures favorable or unfavorable for drug-likeness, offering extra valuable information for getting high quality lead compounds in the early stage of the drug design/discovery process. Finally, the best Bayesian classifier was used to predict the drug-likeness of 33,961 compounds in TCMCD. Our calculations show that 59.37% of the molecules in TCMCD were identified as drug-like molecules, indicating that traditional Chinese medicines (TCMs) are therefore an excellent source of drug-like molecules. Furthermore, the important structural fingerprints in TCMCD were detected and analyzed. Considering that the pharmacology of TCMCD and MDDR (MDL Drug Data Report) was linked by the important common structural features, the potential pharmacology of the compounds in TCMCD may therefore be annotated by these important structural signatures identified from Bayesian analysis, which may be valuable to promote the development of TCMs.
منابع مشابه
Computational investigation of ginsenoside F1 from Panax ginseng Meyer as p38 MAP Kinase Inhibitor: Molecular docking and dynamics simulations, ADMET analysis, and drug likeness prediction.
Ginsenoside F1 is a biologically active compound identified potential from Korean Panax ginseng Meyer. In the present study, the potential targets of ginsenoside F1 were investigated by computational target fishing approaches including ADMET prediction, biological activity prediction from chemical structure, molecular docking, and molecular dynamics methods. Results were suggested to express th...
متن کاملDrug Discovery Using Support Vector Machines. The Case Studies of Drug-likeness, Agrochemical-likeness, and Enzyme Inhibition Predictions
Support Vector Machines (SVM) is a powerful classification and regression tool that is becoming increasingly popular in various machine learning applications. We tested the ability of SVM, in comparison with well-known neural network techniques, to predict drug-likeness and agrochemical-likeness for large compound collections. For both kinds of data, SVM outperforms various neural networks usin...
متن کاملDrug-likeness analysis of traditional Chinese medicines: 1. property distributions of drug-like compounds, non-drug-like compounds and natural compounds from traditional Chinese medicines
UNLABELLED BACKGROUND In this work, we analyzed and compared the distribution profiles of a wide variety of molecular properties for three compound classes: drug-like compounds in MDL Drug Data Report (MDDR), non-drug-like compounds in Available Chemical Directory (ACD), and natural compounds in Traditional Chinese Medicine Compound Database (TCMCD). RESULTS The comparison of the property ...
متن کاملTCMSP: a database of systems pharmacology for drug discovery from herbal medicines
BACKGROUND Modern medicine often clashes with traditional medicine such as Chinese herbal medicine because of the little understanding of the underlying mechanisms of action of the herbs. In an effort to promote integration of both sides and to accelerate the drug discovery from herbal medicines, an efficient systems pharmacology platform that represents ideal information convergence of pharmac...
متن کاملAssessment of "drug-likeness" of a small library of natural products using chemoinformatics
Even though natural products has an excellent record as a source for new drugs, the advent of ultrahigh-throughput screening and large-scale combinatorial synthetic methods, has caused a decline in the use of natural products research in the pharmaceutical industry. This is due to the efficiency in generating and screening a high number of synthetic combinatorial compounds; whereas traditional ...
متن کاملذخیره در منابع من
با ذخیره ی این منبع در منابع من، دسترسی به آن را برای استفاده های بعدی آسان تر کنید
برای دانلود متن کامل این مقاله و بیش از 32 میلیون مقاله دیگر ابتدا ثبت نام کنید
ثبت ناماگر عضو سایت هستید لطفا وارد حساب کاربری خود شوید
ورودعنوان ژورنال:
- Molecular pharmaceutics
دوره 9 10 شماره
صفحات -
تاریخ انتشار 2012